For -Lipschitz function where is the starting radius,
If we run GD for iterations then
If we run GD for iterations with step-size then
(Projected) gradient descent returns with after iterations.
Proof
Claim 1: For all ,
Claim 1(a): For all ,
Claim 1 follows from Claim 1(a) by definition of convexity.
#incomplete
See gradient descent, convergence
References:
- https://www.stat.cmu.edu/~ryantibs/convexopt-F13/scribes/lec6.pdf